Skip to content

Say what the runs say about the weak model's skills delta - #85

Draft
leggetter wants to merge 1 commit into
mainfrom
feature/triage-open-issues-a-875
Draft

leggetter wants to merge 1 commit into
mainfrom
feature/triage-open-issues-a-875

Conversation

@leggetter

Copy link
Copy Markdown
Collaborator

The task

Triage every open issue on this repository: check each is clear, check none has
already been addressed, and propose a roadmap of refine / implement / close. The
triage itself is reported back to the requester — per AGENTS.md, work lives in
GitHub Issues and order lives in #24, not in a markdown file that goes stale.

This PR is the one thing the triage found wrong in the repository rather than
on the board.

What changed

The Status section of AGENTS.md, three claims and one recipe.

1. The recompute snippet counted snapshot files, not executions. A snapshot
republishes the cells a run did not re-measure (#60), so iterating
results/runs/*.json counts one execution many times. The 17 August pair appears
in six snapshot files and reads −2 in all six — which is exactly where the
paragraph's "eight of the nine later runs read −2" came from. The replacement
keys each cell by runId (falling back to ranAt) before counting.

2. "The weak model is two worse with skills than without" is not supported by
any run since August.
Keyed per execution, the five full-suite pairwise runs of
that pair read:

2026-08-13     10/15  13/15  -3      <- credit-outage day
2026-08-17      9/15  11/15  -2
33484972784    14/19  15/19  -1      <- 1 September
34090396233    16/19  15/19  +1      <- 7 September
34813087863    18/19  14/19  +4      <- 14 September

Pooled over the three September executions: +skills 48/57 against 44/57, with
all of the gain on Outpost (14/15 against 10/15) and Event Gateway level at 34/42
in both arms. Those are #2's closing figures, reproduced here from
results/runs/ rather than quoted. #2 was closed as not planned on 22 September
on that basis, and this file kept the old sign for a further two days — and the
old figure for six weeks.

3. Two smaller figures in the same section. "Nine of nineteen scenarios
discriminate" is seven on the published snapshot. And the Claude and GPT-5.6
skills deltas readable from results/latest.json are not deltas: it is a merged
snapshot whose frontier -no-skills arms were executed on 1 September and whose
+skills twins on 14 September, so subtracting them compares two instruments a
fortnight apart. Only the weak pair runs weekly, so only the weak pair is a clean
comparison in that file.

The methodology rule that followed — state the scenario set and the model with
any delta — is kept. What is dropped is its worked claim that no run supports the
sign flipping, because three now do.

What I verified, and how

Every figure above was recomputed from results/runs/ and results/latest.json
in this worktree, not taken from an issue. The snippet as it now appears in the
file was extracted from the markdown and run verbatim; the output in this
description is that run.

pnpm install --frozen-lockfile   ok
pnpm format:check                biome: 148 files, no fixes; prettier: all matched
pnpm typecheck                   framework + web, clean
pnpm check                       20 test files, 227 tests, all passed
pnpm -r build                     done

Nothing was already failing. pnpm install needs /opt/homebrew/bin ahead of
the asdf shims on PATH — the repo has no .tool-versions and the shim has no
pnpm version set. Unrelated to this change; already noted on #82.

No tests

The change is prose and a shell snippet in AGENTS.md. Nothing here is imported
by anything. The snippet was verified by executing it out of the file, which is
the only check available to it.

Deliberately not done

  • No issue was closed, commented on, filed or relabelled. The triage's
    proposals — which issues to close, which to refine, and the order to work in —
    went back to the requester. Closing is a human decision and Roadmap: the order we are working in, and why #24 is the
    requester's to update.
  • Roadmap: the order we are working in, and why #24's body carries the same stale −2 and says milestone 2 is measured at
    "+1 for Claude, 0 for GPT-5.6 and −2 for the deliberately weak model". Correcting
    it is an issue edit, not a diff, and is in the proposal.
  • .plans/delivery-plan.md was not audited. It was last brought back to true
    on 29 August (Bring the plan's status, counts and costs back to what is true #69) and may carry the same figure; checking it is a separate pass
    and would have widened this diff past its description.
  • Line 240's release-cadence paragraph was left alone. It cites −1 on
    1 September and +4 on 14 September to make a different point — that movement is
    not a reason to release — and it is accurate. It omits 7 September's +1, which
    strengthens rather than weakens it.

Observations, not changed here

🤖 Generated with Claude Code

The Status section read "the weak model is two worse with skills than
without" and offered a recompute snippet as the corrective. Both were
wrong, and the snippet is why: it iterates snapshot files rather than
executions, so the 17 August pair appears in six of them and reads -2 in
all six. That is where "eight of the nine later runs read -2" came from.

Keyed per execution, the five full-suite pairwise runs read -3, -2, -1,
+1, +4. #2 closed on 22 September on that basis: pooled across the three
September executions the weak model is +skills 48/57 against 44/57, all
of the gain on Outpost, Event Gateway level at 34/42 in both arms.

Also corrects two other figures in the same section against
results/latest.json: seven of nineteen scenarios are failed by at least
one experiment, not nine, and the frontier deltas in that file are not
deltas at all -- it is a merged snapshot whose -no-skills arms ran on
1 September and whose +skills arms ran on 14 September.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant